Papers with multi-hop reasoning
Copied to clipboard
| Challenge: | Recent advances in reading comprehension have resulted in models that surpass human performance when the answer is contained in a single, continuous passage of text. |
| Approach: | They propose a document-structured message passing architecture for the identification of supporting facts over a graph-structure based representation of text. |
| Outcome: | The proposed model outperforms a baseline reading comprehension test on raw text and shows that it is relevant for multi-hop reasoning. |
Copied to clipboard
| Challenge: | Detailed extended analyses of all submitted systems showed large relative improvements in accessing the most challenging multi-hop inference problems, while absolute performance remains low. |
| Approach: | The Shared Task on Multi-Hop Inference for Explanation Regeneration asks participants to regenerate detailed gold explanations for elementary science questions by selecting facts from a knowledge base of semi-structured tables. |
| Outcome: | The top-performing system achieved a mean average precision of 0.56 . the task combines facts from a knowledge base and supervised training data . |
Copied to clipboard
| Challenge: | Existing open-domain question answering systems only select one source to generate answer or conduct reasoning on structured information. |
| Approach: | They propose a Document-Entity Heterogeneous Graph Network to integrate different sources of information and conduct reasoning on heterogeneous information. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on a HybirdQA dataset. |
Copied to clipboard
| Challenge: | Existing approaches build explanations considering each question in isolation, but new approach leverages explanatory patterns emerging in scientific explanations. |
| Approach: | They propose a framework for reconstructing multi-hop explanations in science Question Answering . they integrate lexical relevance with the notion of unification power to rank atomic facts . |
| Outcome: | The proposed method achieves results competitive with Transformers, but is faster and scalable to large explanatory corpora. |
Copied to clipboard
| Challenge: | Existing studies have shown that large language models can handle knowledge with varying familiarity. |
| Approach: | They propose a benchmark to evaluate multi-hop question answering on new and tail knowledge . they use RAG to integrate external knowledge into large language models . |
| Outcome: | The proposed benchmark evaluates the multi-hop reasoning ability of large language models . it primarily evaluates their ability to handle knowledge with different levels of familiarity . |
Copied to clipboard
| Challenge: | Existing questions that explicitly describe the process for deriving the answer are often implicit. |
| Approach: | They propose a question answering benchmark where the required reasoning steps are implicit in the question and should be inferred using a strategy. |
| Outcome: | The proposed model is short, topic-diverse, and covers a wide range of strategies. |
Copied to clipboard
| Challenge: | Existing datasets for outdoor Vision-and-Language Navigation (VLN) tasks do not include verbal instructions for communicating with mobility. |
| Approach: | They propose a dataset for gesture-guided outdoor VLN instructions with demonstrative expressions that incorporates gestures and linguistic commands. |
| Outcome: | The proposed datasets are compared against existing datasets and analysed in detail. |
Copied to clipboard
| Challenge: | a framework for domain-specific language agents is being developed for industrial automation . a novel approach to adapting these systems to domain-based applications poses new challenges . |
| Approach: | They propose a framework for deploying domain-specific language agents that can query industrial sensor data using natural language. |
| Outcome: | The proposed framework outperforms standard prompting baselines across multiple LLMs including smaller models. |
Copied to clipboard
| Challenge: | Existing hybrid question answering systems use a "prompt-and-pray" paradigm . context size limitations limit ability of many transformer-based LLMs to fit into a given prompt . |
| Approach: | They propose a superset of SQLite to act as a unified dialect for orchestrating reasoning across unstructured and structured data. |
| Outcome: | The proposed framework scales to massive datasets and improves performance while using 35% fewer tokens. |
Copied to clipboard
| Challenge: | Traditional KGQA assumes a closed world where answers must exist in the KG, limiting real-world applicability. |
| Approach: | They propose a system that combines a pre-trained GNN and an LLM for open-world QA. |
| Outcome: | The proposed system outperforms existing LLM–GNN systems on standard benchmarks and GLOW-BENCH, achieving up to 53.3% and an average 38% improvement. |
Copied to clipboard
| Challenge: | GraphRAG integrates structured knowledge graphs into question answering . high-quality triple extraction is critical, but lacks granularity and topical coherence . large language models suffer from inherent limitations in their internalized knowledge . |
| Approach: | They evaluate module-level design choices in GraphRAG for retrieval-augmented generation . they find that triple extraction is critical for accurate and comprehensive retrieval . |
| Outcome: | The proposed framework outperforms other retrieval-augmented generation frameworks in accuracy and efficiency. |
Copied to clipboard
| Challenge: | Existing knowledge-based visual question answering tasks require weak supervision and no visual knowledge. |
| Approach: | They propose a model which encodes high-level semantics of a question and a knowledge base and learns high order associations between them. |
| Outcome: | The proposed model encodes high-level semantics of a question and a knowledge base, and learns high order associations between them. |
Copied to clipboard
| Challenge: | Existing methods struggle to conduct deep searches and retrieve all necessary evidence. |
| Approach: | They propose a benchmark for evaluating deep search, a retrieval-augmented generation that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources. |
| Outcome: | The proposed benchmarks show that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on the benchmark. |
Copied to clipboard
| Challenge: | Existing models trained on poor quality data have shown strong performance in language modeling and some downstream benchmarks. |
| Approach: | They evaluate kNN-LMs on a diverse set of tasks and evaluate their performance. |
| Outcome: | The proposed extension could improve on a variety of tasks, but it fails to perform on reasoning tasks that require integrating multiple pieces of information. |
Copied to clipboard
| Challenge: | Existing MAS frameworks often require manual workflow configuration and lack native support for dynamic evolution and performance optimization. |
| Approach: | They propose an open-source platform that automates generation, execution, and evolutionary optimization of multi-agent workflows. |
| Outcome: | The proposed platform automates generation, execution, and evolutionary optimization of multi-agent workflows. |
Copied to clipboard
| Challenge: | Existing language models are inadequate in reasoning, according to studies . a new reasoning pre-training paradigm is based on pretraining language models with programs . |
| Approach: | They propose a reasoning pre-training paradigm that empowers language models to harvest reasoning knowledge possessed by program executors. |
| Outcome: | The proposed reasoning pre-training paradigm can boost models' reasoning skills . it can be instantiated by different kinds of program executors and run on a single database . |
Copied to clipboard
| Challenge: | Experimental evaluations on counseling dialogue dataset, POEM validate MENDER’s efficacy in generating coherent, knowledge-grounded responses. |
| Approach: | They propose a multi-hop commonsensE and domaiN-specific Chain-of-Thought reasoning framework that integrates commonsense and domain knowledge via multi-hopping reasoning over the dialogue context. |
| Outcome: | Experimental evaluations on counseling dialogue dataset validate MENDER’s efficacy in generating coherent, empathetic, knowledge-grounded responses. |
Copied to clipboard
| Challenge: | Existing methods for multi-hop relation reasoning require limited data for each query relation, resulting in limited interpretation. |
| Approach: | They propose a few-shot multi-hop relation learning model that uses reinforcement learning to model sequential steps of multi-hopping reasoning and performs heterogeneous structure encoding and knowledge-aware search space pruning. |
| Outcome: | Empirical results show that the proposed model outperforms state-of-the-art models over few-shot relations. |
Copied to clipboard
| Challenge: | Existing locate-and-edit knowledge editing methods suffer from two limitations: they are infeasible for large scale KE in practice and require long run-time. |
| Approach: | They propose to use parametric fine-tuning techniques to update obsolete knowledge and induce new knowledge into LLMs. |
| Outcome: | The proposed methods improve the performance of KE and knowledge update in a temporal dataset with knowledge update and knowledge injection examples. |
Copied to clipboard
| Challenge: | Existing studies suggest key phrase selection is essential for question generation, yet it is difficult to connect disjointed phrases into meaningful questions, especially for long context. |
| Approach: | They propose a QG framework that uses multi-level content planning to generate questions from a given context and an answer. |
| Outcome: | The proposed framework outperforms baselines on two popular QG datasets. |
Copied to clipboard
| Challenge: | Existing multi-hop reading comprehension datasets have reasoning shortcuts that can be used to answer comparison questions without performing multi- hop reasoning. |
| Approach: | They propose a dataset with three probing tasks in addition to the main question . they then evaluate the model's ability to understand date information . |
| Outcome: | The proposed model performs well in date comparison and number subtraction tasks. |
Copied to clipboard
| Challenge: | Existing methods for NER and RE annotation are costly and difficult to scale. |
| Approach: | They propose a semantic stability framework for constructing explainable KGs using NER and RE annotations. |
| Outcome: | The proposed framework supports multi-hop reasoning, triadic SUD–SDOH–SUD mediation patterns, and feedback loop analysis. |
Copied to clipboard
| Challenge: | Existing methods that decompose multi-hop questions into single hop sub-questions are difficult to implement. |
| Approach: | They propose to use random-walks to guide pre-trained language models to map multi-hop questions to random-walked paths that lead to the answer. |
| Outcome: | The proposed methods improve on two T5 LMs. |
Copied to clipboard
| Challenge: | Multi-hop reasoning requires a chain of facts to reflect the reasoning behind the answer. |
| Approach: | They propose an inference-guided prompting approach that performs well in natural language questions . they propose a neuro-symbolic approach to reasoning using large language models . |
| Outcome: | The proposed model outperforms all prompting strategies and fine-tunes LLMs trained specifically for proof generation. |
Copied to clipboard
| Challenge: | Existing hallucination detection frameworks for RAGs lack robustness and performance . a compact model may lose track of precise information in retrieved segments or misinterpret a document's entailment score. |
| Approach: | They propose a lightweight, modular framework for hallucination detection in RAG systems . they capture logical relationships among retrieved documents within the vector space . |
| Outcome: | The proposed framework improves hallucination detection in RAG systems without complex architectures or pre-training on datasets. |
Copied to clipboard
| Challenge: | Existing agentic approaches for Knowledge Graph-based Retrieval-Augmented Generation fail to generalize to real-world enterprise Knowledge graphs (KGs) dense, schema-driven, and operationally constrained, requiring a training-free framework. |
| Approach: | They propose a training-free framework that integrates structured planning with controlled iterative reasoning by injecting schema-conditioned structural priors and enforcing schemas during multi-hop reasoning. |
| Outcome: | The proposed framework significantly improves on a real-world enterprise-oriented benchmark constructed from a Configuration Management DataBase (CMDB). |
Copied to clipboard
| Challenge: | Existing models extract evidence in both sentences and table cells from Wikipedia dumps, ignoring potential connections between them. |
| Approach: | They propose a model which uses a mixed evidence graph to extract the evidence in both formats without manually designed conversion rules. |
| Outcome: | The proposed model outperforms existing models and improves the verification step. |
Copied to clipboard
| Challenge: | Existing models for commonsense reasoning are limited by their limited set of facts, rendering them unfit for reasoning over new unseen situations and events. |
| Approach: | They propose a neural-symbolic reasoner which can combine commonsense facts with large-scale dynamic CKGs to draw conclusions about ordinary situations. |
| Outcome: | The proposed model outperforms the state-of-the-art models on the task of link prediction on CKGs. |
Copied to clipboard
| Challenge: | In implicit sentiment analysis, the opinion cues come in an implicit and obscure manner. |
| Approach: | They propose a three-step prompting principle for THOR to step-by-step induce the implicit aspect, opinion and finally the sentiment polarity. |
| Outcome: | The proposed framework pushes the state-of-the-art (SoTA) by over 6% F1 on supervised setup and more strikingly, boosts the SoTA by over 50% F1 with THOR+GPT3. |
Copied to clipboard
| Challenge: | Existing approaches for multi-hop reasoning are lacking for local graph reasoning . existing approaches neglect local semantic structures in utterances . |
| Approach: | They propose a question-aware global-to-local graph reasoning approach that expands the canonical Interlocutor-Utterance graph by introducing a query node. |
| Outcome: | The proposed approach outperforms existing methods on Molweni and FriendsQA. |
Copied to clipboard
| Challenge: | Recent years have seen the proliferation of disinformation and fake news online. |
| Approach: | They propose to model the context of a political debate and the contexts of the document describing the fact-checked claim. |
| Outcome: | The proposed model improves on the state-of-the-art model by modeling the context of the claim . the experimental results show that the model can provide 10+ points of improvement over the state of the art model . |
Copied to clipboard
| Challenge: | Recent advances in long-context modeling have enhanced language models for complex tasks, but they struggle with multi-hop reasoning and noisy contexts. |
| Approach: | They propose an approach that prompts LMs to supply attributions for each assertion during reasoning. |
| Outcome: | The proposed model achieves competitive performance on multi-hop reasoning benchmarks, closely paralleling proprietary LMs such as ChatGPT and Claude-instant. |
Copied to clipboard
| Challenge: | Existing reasoning models suffer from hallucinations and unfaithfulness, whereas general LLMs perform suboptimal on complex tasks. |
| Approach: | They propose a structure analysis method that helps LLMs better understand the question structure and guide the problem-solving process. |
| Outcome: | The proposed method improves zero-shot performance on knowledge-intensive and mathematical tasks while demonstrating strong robustness against corrupted reasoning paths. |
Copied to clipboard
| Challenge: | Existing methods for document-level relation extraction capture non-local interactions but are not able to capture rich non-linguistic interactions. |
| Approach: | They propose a document-level relation extraction model that empowers relational reasoning across sentences by automatically inducing the latent document- level graph. |
| Outcome: | The proposed model achieves an F1 score of 59.05 on a large-scale document-level dataset (DocRED), significantly improving over the previous results. |
Copied to clipboard
| Challenge: | Existing methods for multi-hop reasoning ignore grounding on supporting facts of each step, which tends to generate inaccurate decompositions. |
| Approach: | They propose an interpretable stepwise reasoning framework that incorporates supporting sentences and questions at each intermediate step and utilizes the inference of the current hop for the next until reasoning out the final result. |
| Outcome: | The proposed model can boost performance and yield a better interpretable reasoning process without decomposition supervision. |
Copied to clipboard
| Challenge: | Existing methods focus on extracting relations from single sentence . document-level relation extraction requires a comprehension of the whole document . |
| Approach: | They propose a graph-based model with Dual-tier Heterogeneous Graph (DHG) for document-level relation extraction. |
| Outcome: | The proposed model achieves state-of-the-art performance on two widely used datasets. |
Copied to clipboard
| Challenge: | State-of-the-art Large Language Models (LLMs) are accredited with a number of different capabilities, including reading comprehension, mathematical and reasoning skills, and possessing scientific knowledge. |
| Approach: | They propose a benchmark to generate seemingly plausible multi-hop reasoning chains that ultimately lead to incorrect answers. |
| Outcome: | The proposed model circumvents the reasoning requirement but in subtle ways . it shows that it is more difficult to generate plausible alternatives . |
Copied to clipboard
| Challenge: | Existing methods for open-domain table question answering require retraining or fine-tuning on new datasets. |
| Approach: | They propose a zero-shot, cascaded retrieval approach that uses a sparse retrieval model to filter a subset of candidates before applying more expensive dense models as re-rankers. |
| Outcome: | The proposed method outperforms state-of-the-art retrieval models on the NQ-Tables dataset. |
Copied to clipboard
| Challenge: | Recent years have witnessed interest in Temporal Question Answering over Knowledge Graphs (TKGQA) but these methods are highly engineered and do not automatically discover relevant parts of the KG during multi-hop reasoning. |
| Approach: | They propose a scheme to modulate the messages passed through a KG edge during convolution based on the relevance of its associated period to the question. |
| Outcome: | The proposed system outperforms state-of-the-art models on a recent challenging dataset for multi-hop complex temporal QA called TimeQuestions. |
Copied to clipboard
| Challenge: | KG-MulQA extracts QA pairs at multiple complexity levels along three key dimensions: multi-hop retrieval, set operations, and answer plurality. |
| Approach: | They propose a framework that extracts QA pairs at multiple complexity levels along three key dimensions: multi-hop retrieval, set operations, and answer plurality. |
| Outcome: | The framework extracts QA pairs at multiple complexity levels along key dimensions . it enables fine-grained assessment of model performance across controlled difficulty levels. |
Copied to clipboard
| Challenge: | Generative question answering (QA) models generate answers to complex questions, but their mechanism for doing so is still poorly understood. |
| Approach: | They decompose multi-hop questions into multiple corresponding single-hop question chains and find marked inconsistency in QA models’ answers on these pairs of ostensibly identical question chains. |
| Outcome: | The proposed models lack zero-shot multi-hop reasoning ability when trained on single-hop questions and on logical forms. |
Copied to clipboard
| Challenge: | Existing approaches to answer natural language questions on knowledge graphs (KGQA) use large-scale entity-related text corpus or knowledge graph embeddings as auxiliary information to facilitate answer selection. |
| Approach: | They propose to integrate explicit textual information and implicit KG structural features of relation paths into a novel rotate-and-scale entity link prediction framework. |
| Outcome: | The proposed method is superior to existing methods on three KGQA datasets and shows that it can be used to identify answer entities. |
Copied to clipboard
| Challenge: | Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes, but lacks standardized evaluation suites for non-English languages. |
| Approach: | They propose a novel benchmark that includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. |
| Outcome: | The proposed benchmark includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. |
Copied to clipboard
| Challenge: | Extensive experimental results on several popular logical benchmarks (ProofWriter, PrOntoQA, PrONtoQA-OOD, and FOLIO) and mathematical benchmark (DI-GSM) show that COP significantly outperforms previous state-of-the-art methods. |
| Approach: | They propose a reasoning approach called Concise and Organized Perception (COP) that carefully analyzes the given statements to identify the most pertinent information while eliminating redundancy efficiently. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on several popular logical benchmarks and mathematical benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable performances in general domains and are now extending into the expert domain of law. |
| Approach: | They propose a Korean Benchmark for Legal EXplainable QA (KoBLEX) that evaluates provision-grounded, multi-hop legal reasoning. |
| Outcome: | The proposed method outperforms baselines and shows a high correlation with human judgments. |
Copied to clipboard
| Challenge: | Traditional supervised QG methods rely on tokenlevel alignment with fixed gold labels struggle to capture diverse valid question formulations. |
| Approach: | They propose a model-agnostic framework that integrates multimodal inputs with a multi-decoder architecture to optimize for multiple labels per sample. |
| Outcome: | The proposed framework improves fluency, reasoning depth, and relevance of visual questions. |
Copied to clipboard
| Challenge: | Question Answering over Knowledge Graph (KGQA) aims to find answer entities for natural language questions based on knowledge graphs. |
| Approach: | They propose a subgraph-aware self-attention mechanism to imitate the graph neural network (GNN) based module to perform multi-hop reasoning on KG. |
| Outcome: | The proposed method surpasses state-of-the-art models by a large margin even with fewer updated parameters and less training data. |
Copied to clipboard
| Challenge: | Existing QA models rely on shortcuts to provide the true answer, referred to as disconnected reasoning problem. |
| Approach: | They propose a causal-effect approach that exploits true multi-hop reasoning instead of shortcuts. |
| Outcome: | The proposed method achieves 5.8% higher points of its Supps score on hotpotQA through true multihop reasoning. |
Copied to clipboard
| Challenge: | Naive Retrieval-Augmented Generation (RAG) focuses on individual documents during retrieval and is not suitable for networked documents. |
| Approach: | They propose a novel divide-and-conquer strategy that retrieves optimal subgraph structure in linear time. |
| Outcome: | The proposed approach outperforms current state-of-the-art methods on graph reasoning benchmarks. |
Copied to clipboard
| Challenge: | Current evaluations of RAG systems overlook structural complexity and multi-step reasoning . GRADE model enables fine-grained analysis of Ragging performance . |
| Approach: | They propose a framework that models retrieval difficulty along two orthogonal dimensions . they extract knowledge graphs and augment them through semantic clustering to recover missing links . |
| Outcome: | The proposed framework models retrieval difficulty along two orthogonal dimensions . error rates correlate with the framework, and it validates its diagnostic utility. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been significantly improved by instruction fine-tuning, but still lack transparency and the ability to utilize up-to-date knowledge and information. |
| Approach: | They propose a search-augmented instruction learning model which grounds the language generation and instruction following abilities on complex search results generated by in-house and external search engines. |
| Outcome: | The proposed model outperforms plain LLMs on zero-shot language tasks and can generate both natural and programming languages following natural language guidance and requests. |
Copied to clipboard
| Challenge: | Empirical evaluation shows our model to outperform the single-hop question generation models on both automatic evaluation metrics such as BLEU, METEOR, and ROUGE and human evaluation metrics for quality and coverage of the generated questions. |
| Approach: | They propose a question-aware reward function to maximize the utilization of supporting facts in the context. |
| Outcome: | The proposed model outperforms single-hop neural question generation models on automatic evaluation metrics and human evaluation metrics for quality and coverage of the generated questions. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown potential in reasoning over structured environments, e.g., knowledge graphs and tables. |
| Approach: | They propose a framework that allows LLMs to efficiently and faithfully reason over structured environments. |
| Outcome: | The proposed framework surpasses state-of-the-art fine-tuned methods on three KGQA and two TableQA datasets and surpasse CWQ and WTQ methods. |
Copied to clipboard
| Challenge: | Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI) data and synthetic data. |
| Approach: | They propose a method that leverages multi-hop reasoning on context graphs extracted from documents to generate complex multi-level claims without relying on LLMs to decide data labels. |
| Outcome: | The proposed model outperforms GPT-4-o on the LLM-Aggrefact benchmark with much smaller model size. |
Copied to clipboard
| Challenge: | Currently, one-step retrieve-and-read question answering systems cannot answer such questions because they rarely contain retrievable clues about the missing entity. |
| Approach: | They propose a multi-step approach to retrieve relevant content with the question, then reading the paragraphs returned by the information retrieval component to arrive at the final answer. |
| Outcome: | The proposed model outperforms the best previously published model despite not using pretrained language models such as BERT. |
Copied to clipboard
| Challenge: | Transformers are used to solve multi-hop question answering tasks that require reasoning over multiple parts of a long document. |
| Approach: | They propose a method that collects relevant information over the entire document and then combines it with local context to solve a multi-hop question answering task. |
| Outcome: | The proposed method improves on three MHQA datasets compared to the baseline model. |
Copied to clipboard
| Challenge: | In multi-hop question answering, models need to connect multiple pieces of evidence scattered in a long context to answer the question. |
| Approach: | They propose to use a control unit that dynamically attends to the question at different reasoning hops to guide the model's multi-hop reasoning. |
| Outcome: | The proposed model outperforms baseline models but is limited on adversarial test. |
Copied to clipboard
| Challenge: | Existing approaches to multi-hop reading comprehension do not include multiple sentences or passages. |
| Approach: | They propose a path-based reasoning approach for a multi-hop reading comprehension task . they propose to extract paths from text and compose them to encode them . |
| Outcome: | The proposed model outperforms previous models on the multi-hop Wikihop dataset and can be generalized to the OpenBookQA dataset. |
Copied to clipboard
| Challenge: | Extensive experiments across seven long-context tasks demonstrate that AgenticLU significantly outperforms state-of-the-art prompting methods and specialized long-consumer LLMs. |
| Approach: | They propose a framework to enhance an LLM's understanding of long-context questions by integrating targeted self-clarification with contextual grounding within an agentic workflow. |
| Outcome: | The proposed framework outperforms state-of-the-art prompting methods and specialized long-context LLMs in seven long-constitut tasks. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue systems generate accurate and informative responses, but they are prone to hallucination problems. |
| Approach: | They propose a method to generate hallucinated responses using knowledge graphs . they propose local knowledge grounding to combine textual embeddings with corresponding KG embeddments . a global knowledge ground technique is also proposed to equip RHO with multi-hop reasoning abilities . |
| Outcome: | The proposed approach outperforms state-of-the-art methods on automatic and human evaluation by a large margin. |
Copied to clipboard
| Challenge: | Knowledge graphs are incomplete with many facts missing, causing performance bottlenecks in many applications. |
| Approach: | They propose a general multi-hop reasoning task that can be formulated as a search process and can be extended to long-distance reasoning scenarios. |
| Outcome: | The proposed model improves on baselines in short and long distance reasoning scenarios. |
Copied to clipboard
| Challenge: | Recent advances in Graph-based RAG (GRAG) frameworks focus on knowledge graphs for cross-lingual retrieval. |
| Approach: | They propose a new GRAG framework for cross-lingual question answering . MaGiX constructs a multi-granular cross-linguistic knowledge graph using fine-grained attribute descriptions and cross-synonym edges. |
| Outcome: | The proposed framework outperforms prior GRAG systems in retrieval accuracy and generation quality. |
Copied to clipboard
| Challenge: | Existing methods to answer complex questions rely on decomposition of complex questions into sub-questions . Existing approaches to decompose complex questions are limited by the original question . |
| Approach: | They propose a question decomposition approach to decompose semantically clear questions . they use the decomposed sub-questions to select relevant patterns as auxiliary information . |
| Outcome: | The proposed method achieves state-of-the-art performance on multiple datasets. |
Copied to clipboard
| Challenge: | Existing models fail to answer a large portion of sub-questions . Existing systems have achieved super-human performance . |
| Approach: | They propose to use a neural decomposition model to generate sub-questions for a multi-hop question and extract the corresponding sub-answers. |
| Outcome: | The proposed model is based on a hotpotQA dataset with a multi-hop question and sub-answers. |
Copied to clipboard
| Challenge: | Knowledge Base Question Answering (KBQA) systems have limited generalizability across knowledge bases and multiple reasoning types. |
| Approach: | They propose a modular approach for KBQA that is built on a framework adaptable to multiple knowledge bases and reasoning types. |
| Outcome: | The proposed approach is generalized across multiple knowledge bases and reasoning types. |
Copied to clipboard
| Challenge: | Existing approaches to multihop reasoning fail to address the problem of spurious paths . existing approaches neglect the internal semantic consistency of the reward function . |
| Approach: | They propose a framework that incorporates semantic consistency into the reward function to guide multi-hop reasoning. |
| Outcome: | The proposed framework outperforms baseline methods and facilitates more interpretable reasoning paths. |
Copied to clipboard
| Challenge: | Graph-based RAG systems have been promising for enabling multi-hop reasoning . but when knowledge graphs are constructed from unstructured documents, they suffer from fragmentation . |
| Approach: | They propose a framework to reconstruct and enrich fragmented knowledge graphs . they propose three core components: Graph Reorganization, Perspective Expansion, and Query-aware Reranking. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks on four benchmarks . it achieves over 80% diversity win rate and enables multi-hop reasoning . |
Copied to clipboard
| Challenge: | Existing approaches to question answering on knowledge graphs are based on a modularized sequential approach where errors in one module lead to the accumulation of errors in downstream modules. |
| Approach: | They propose a multi-task BERT based Neural Machine Translation model to address these challenges. |
| Outcome: | The proposed model can answer questions over a knowledge graph on one publicly available and one proprietary dataset. |
Copied to clipboard
| Challenge: | Existing approaches to augment language models with external knowledge but they are limited by static nature of pre-training data. |
| Approach: | They propose a lightweight approach that compresses retrieved documents into highly dense textual summaries to integrate into in-context RAG. |
| Outcome: | The proposed approach reduces latency and costs while achieving high performance in open-domain questions. |
Copied to clipboard
| Challenge: | Existing studies on knowledge editing focus on monolingual scenarios, neglecting the complexities presented by multilingual contexts and multi-hop reasoning. |
| Approach: | They propose a benchmark to evaluate the adaptability of multilingual knowledge editing methods. |
| Outcome: | The proposed benchmark evaluates the adaptability of multilingual knowledge editing methods across five languages. |
Copied to clipboard
| Challenge: | Existing Key-value Memory Neural Networks are effective for shallow reasoning over documents . but extending them to Knowledge Based Question Answering is not trivial . |
| Approach: | They propose a mechanism to enable conventional KV-MemNNs models to perform interpretable reasoning for complex questions. |
| Outcome: | The proposed solution provides better reasoning abilities on complex questions and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Retrieval Augmented Generation (RAG) is a non-parametric approach for large language models. |
| Approach: | They propose a framework that shifts from triples to context-rich propositions and introduces an efficient, LLM-free online beam search over proposition paths to discover multi-step reasoning chains. |
| Outcome: | The proposed framework achieves state-of-the-art zero-shot Recall@5 and F1 scores on 2Wiki, HotpotQA, and MuSiQue. |
Copied to clipboard
| Challenge: | Existing reasoning path retrieval methods lack a global structural perspective. |
| Approach: | They propose a framework that reframes multi-hop reasoning as a schema-guided graph search task. |
| Outcome: | The proposed framework improves accuracy and evidence completeness of multi-hop reasoning graph retrieval. |
Copied to clipboard
| Challenge: | Existing approaches to interpret task-oriented dialogue systems employ an implicit reasoning strategy that makes the model predictions uninterpretable to humans. |
| Approach: | They propose a neuro-symbolic approach that performs explicit reasoning that justifies model decisions by reasoning chains. |
| Outcome: | The proposed approach achieves better results and introduces an interpretable decision process. |
Copied to clipboard
| Challenge: | a human-like chatbot requires commonsense reasoning to comprehend and respond to information . however, identifying and aggregating key evidence within a single hop is a challenge . a knowledge distillation framework is proposed that leverages LLMs as unreliable teachers . |
| Approach: | They propose a framework that leverages large language models as unreliable teachers to facilitate multi-hop reasoning over a dialogue context. |
| Outcome: | The proposed framework leverages LLMs as unreliable teachers and selectively distills consistent and helpful rationales via alignment filters. |
Copied to clipboard
| Challenge: | et al. : evidence retrieval is highly dependent on partial, incorrect or no supporting knowledge. |
| Approach: | They propose a method that retrieves and reranks evidence facts jointly . they propose to account for links between sentences and coverage with the given query . |
| Outcome: | The proposed approach achieves state-of-the-art evidence retrieval performance on two multi-hop question answering datasets. |
Copied to clipboard
| Challenge: | Existing legal NLP benchmarks focus on single-clause tasks, such as ContractNLI and CUAD. |
| Approach: | They propose a framework that models cross-clause dependencies through structured clause graphs by extracting deontic-temporal entities from clauses and constructs typed relationship graphs capturing definitional dependencies, exception hierarchies, and temporal sequences. |
| Outcome: | The proposed framework extracts deontic-temporal entities from clauses and constructs typed relationship graphs capturing definitional dependencies, exception hierarchies, and temporal sequences. |
Copied to clipboard
| Challenge: | Existing datasets that explicitly focus on multi-hop reasoning are lacking in learning multi-tasking. |
| Approach: | They propose to use sentence-factored models to solve multi-hop question answering tasks . they find spurious correlations in unmasked versions of WikiHop and HotpotQA . |
| Outcome: | The proposed datasets are used to test models on multi-hop question answering tasks. |
Copied to clipboard
| Challenge: | Existing knowledge editing methods overlook interplay with pre-existing knowledge, leading to inconsistent edit propagation. |
| Approach: | stepKE integrates edited and existing knowledge for coherent multi-hop reasoning . stepKE decomposes multi-step questions into sequential single-hop sub-questions . |
| Outcome: | Experiments show that StepKE generates more accurate and consistent responses than baselines. |
Copied to clipboard
| Challenge: | Existing studies on text-based QG focus on generating SQuAD-style questions. |
| Approach: | They propose a multi-hop question generation model that does context encoding in multiple hops with Graph Convolutional Network and encoder fusion via an Encoder Reasoning Gate. |
| Outcome: | Empirical results show that the proposed model generates fluent questions with high completeness and outperforms baselines on automatic evaluation metrics. |
Copied to clipboard
| Challenge: | a single-hop reasoning model can solve much more of the dataset than previously thought. |
| Approach: | They propose a single-hop BERT-based RC model that achieves 67 F1 . they propose an evaluation setting where humans are not shown all paragraphs . |
| Outcome: | The proposed model achieves 67 F1—comparable to state-of-the-art multi-hop models. |
Copied to clipboard
| Challenge: | Existing approaches struggle to efficiently navigate complex codebases when identifying relevant code snippets. |
| Approach: | They propose a graph-guided agent framework that addresses code localization through a distributed graph-based agent. |
| Outcome: | The proposed framework improves accuracy on real-world benchmarks and can be used to locate code snippets at a cost of 86%. |
Copied to clipboard
| Challenge: | Long-context Document Visual Question Answering (DocVQA) methods struggle with visual semantics or handling finite context windows. |
| Approach: | They propose a new approach to longcontext document visual question answering that transforms retrieval into adaptive evidence chain construction using a Bi-Layered Graph. |
| Outcome: | The proposed approach achieves an average accuracy improvement of 14.07% on M5BookVQA and exhibits robust generalization with a 13.38% gain across four established benchmarks. |
Copied to clipboard
| Challenge: | Neural models, including large language models (LLMs), achieve superior performance on multi-hop question-answering tasks. |
| Approach: | They propose to use the chain-of-thought mechanism to generate both the reasoning chain and the answer. |
| Outcome: | Empirical results show that the proposed framework generates more faithful reasoning chains and significantly improves the QA performance on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing models for multi-hop question answering require multiple pieces of evidence scattered in a given context. |
| Approach: | They propose an interpretable, controller-based self-assembling Neural Modular Network for multi-hop reasoning . their model can softly decompose a multi-step question into multiple single-hop sub-questions . |
| Outcome: | The proposed model improves on the static, single-hop model on regular and adversarial evaluations. |
Copied to clipboard
| Challenge: | Existing datasets that focus on temporal knowledge are limited in size and lack comprehensive coverage of temporal information. |
| Approach: | They introduce a large-scale temporal question-answer-matching dataset . the new taxonomy categorizes questions as attributes, comparisons, and counting questions . |
| Outcome: | The proposed dataset surpasses existing benchmarks in scale and scope. |
Copied to clipboard
| Challenge: | Existing benchmarks for investigating knowledge conflict have notable limitations, including a narrow focus on the question answering setup, heavy reliance on entity substitution techniques, and a restricted range of conflict types. |
| Approach: | They propose a knowledge graph-based framework that generates varied and subtle conflicts between two similar yet distinct contexts while ensuring interpretability through the explicit relational structure of KGs. |
| Outcome: | The proposed framework generates varied and subtle conflicts between two similar yet distinct contexts while ensuring interpretability through the explicit relational structure of KGs. |
Copied to clipboard
| Challenge: | Current models can not ensure the complexity of generated questions, so they may generate shallow questions that can be answered without multi-hop reasoning. |
| Approach: | They propose a controlled framework to generate multi-hop questions that contain key entities in multi- hop reasoning chains and a novel Transformer-based decoder to guarantee that key entities appear in the questions. |
| Outcome: | The proposed model outperforms the state-of-the-art model 25% on HotpotQA. |
Copied to clipboard
| Challenge: | Existing approaches to robustify multi-hop question answering models require expensive annotations. |
| Approach: | They propose a method to supervise answers with right reasoning chains without annotations . they compare answers confidence with and without evidence sentences to generate "pseudo-evidentiality" annotations. |
| Outcome: | The proposed model is accurate and robust in multi-hop reasoning. |
Copied to clipboard
| Challenge: | Existing knowledge graph question answering methods rely on LLM-induced type systems with inconsistent granularity or perform multi-hop reasoning without explicit target-type constraints. |
| Approach: | They propose a type-constrained knowledge graph question answering framework that reasons over a relation-centric ontology graph. |
| Outcome: | The proposed framework achieves state-of-the-art and produces ontology-grounded reasoning chains with substantial Hit@1 gains. |
Copied to clipboard
| Challenge: | Existing retrievers are not perfect and often include irrelevant documents in the retrieved set. |
| Approach: | They propose to construct knowledge-grounded reasoning chains from retrieved documents to integrate supporting evidence into RAG models. |
| Outcome: | The proposed model achieves an average performance improvement of 14.03% on three multi-hop QA datasets. |
Copied to clipboard
| Challenge: | Generating multiple-choice questions (MCQG) for professional exams is challenging due to outdated knowledge, hallucination issues, and prompt sensitivity. |
| Approach: | They propose a framework for converting medical cases into high-quality USMLE-style questions using a self-refine-based framework. |
| Outcome: | The proposed framework improves human expert satisfaction regarding quality and difficulty of medical questions. |
Copied to clipboard
| Challenge: | Existing methods for summarizing source document for non-factoid questions are lacking in factoidic QA. |
| Approach: | They propose a question-driven abstractive summarization method that incorporates multi-hop reasoning into question-based summarizing. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two non-factoid QA datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of ‘Superstition’ is". |
| Approach: | They examine whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of ‘Superstition’ is". |
| Outcome: | The proposed model can latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of ‘Superstition’ is". |
Copied to clipboard
| Challenge: | Existing language model pretraining methods do not capture dependencies or knowledge that span across documents. |
| Approach: | They propose a language model pretraining method that leverages links between documents . they use masked language modeling and document relation prediction to model LMs . |
| Outcome: | The proposed method outperforms existing methods on downstream tasks across two domains. |
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks focus on single-chart tasks, neglecting multi-hop reasoning required to extract and integrate information from multiple charts. |
| Approach: | They propose a benchmark that evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
| Outcome: | The proposed benchmark evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
Copied to clipboard
| Challenge: | Increasing the ability of large language models to perform latent multihop reasoning is crucial for reducing the cost and deployment challenges. |
| Approach: | They propose an interpretability method that traces how logits propagate across layers and positions toward the final prediction. |
| Outcome: | The proposed method improves accuracy on five reasoning datasets. |
Copied to clipboard
| Challenge: | MRC in other languages, including Russian, has not been well-addressed due to the lack of high-quality and large-scale datasets. |
| Approach: | They propose two Russian machine reading comprehension datasets that require reasoning over multiple sentences and commonsense knowledge to infer the answer. |
| Outcome: | The proposed datasets are more complex than the original ones for Russian . the results show that the proposed models are challenging for advanced models . |
Copied to clipboard
| Challenge: | Existing methods for multi-hop reasoning in knowledge base question answering are coarse-grained and may bring information loss. |
| Approach: | They propose a sequential reasoning self-attention mechanism to capture the crucial reasoning information of each hop in a more fine-grained way. |
| Outcome: | The proposed model achieves new state-of-the-art Hits@1 of 76.8% on WebQSP and is also effective when KB is incomplete. |
Copied to clipboard
| Challenge: | Existing multi-hop question answering datasets do not provide a complete explanation for the reasoning process from the question to the answer. |
| Approach: | They propose a multi-hop question answering dataset that uses structured and unstructured data to test reasoning skills. |
| Outcome: | The proposed dataset ensures multi-hop reasoning while being challenging for multi-models. |
Copied to clipboard
| Challenge: | Currently, conversational agents lack commonsense reasoning, preventing them from engaging in rich conversations with humans. |
| Approach: | They propose a commonsense reasoning system that uncovers unstated presumptions from user commands satisfying a general template of if-(state), then-(action), because-(goal) They propose to use a transformer-based generative commons sense knowledge base as its source of background knowledge to extract multi-hop reasoning chains from the neural KB. |
| Outcome: | The proposed model achieves a 35% higher success rate than existing methods with human users. |
Copied to clipboard
| Challenge: | Existing work on this task either evaluates facts in isolation or artificially limits the possible chains of facts, thus limiting multi-hop inference. |
| Approach: | They propose an iterative inference algorithm that decomposes the selection of facts from a corpus autoregressively and conditioning the next iteration on previously selected facts. |
| Outcome: | The proposed method outperforms the previous state-of-the-art in terms of precision, training time and inference efficiency on the WorldTree dataset. |
Copied to clipboard
| Challenge: | Existing approaches to multi-hop question answering emphasize single-step and multi-step iterative decomposition or retrieval, which are susceptible to failure in long-chain reasoning due to the progressive accumulation of erroneous information. |
| Approach: | They propose a Local-tO-Global optimized retrieval method to discover more beneficial information and improve tuplet objective loss. |
| Outcome: | The proposed method outperforms state-of-the-art models and significantly improves multi-hop reasoning. |
Copied to clipboard
| Challenge: | Text-based question answering (TBQA) has been studied extensively in recent years. |
| Approach: | They propose a Dynamically Fused Graph Network to answer questions requiring multiple scattered evidence and reasoning over them. |
| Outcome: | The proposed method achieves competitive results on a public TBQA dataset and produces interpretable reasoning chains. |
Copied to clipboard
| Challenge: | Existing evidence that humans make numerous inferences to understand discourse and text is not fully understood. |
| Approach: | They propose to use textual inference datasets with multi-sentence premises to solve the entailment verification problem. |
| Outcome: | The proposed model outperforms GPT-3.5 and rivals GPL-4 in EV tasks. |
Copied to clipboard
| Challenge: | Existing methods rely on unstructured retrieval or coarse abstractions, which lead to temporal conflicts, brittle reasoning, and limited traceability. |
| Approach: | They propose a unified memory framework that consolidates long-term agent experiences into three interconnected components that combine structured knowledge and evidence to construct compact yet information-dense contexts for reasoning. |
| Outcome: | The proposed framework significantly improves multi-hop and temporal reasoning accuracy while reducing input context length by over 95% compared to long-context baselines. |
Copied to clipboard
| Challenge: | Existing studies have focused on assessing the model’s overall accuracy without evaluating it on different reasoning cases. |
| Approach: | They propose a novel idea to identify and improve multi-modal multi-hop reasoning in VQA by using two new language prompts to find a reasoning path to reach its answer. |
| Outcome: | The proposed model improves multi-modal multi-hop reasoning in visual question answering (VQA) it finds that the proposed model is easy to answer, simply demanding “single-hop” reasoning, whereas only a few questions require “multi-hop.” |
Copied to clipboard
| Challenge: | Relation Extraction (RE) is a task that seeks to identify the relation of entities described according to some context. |
| Approach: | They propose a multi-hop evidence retrieval method based on evidence path mining and ranking to support cross-document relation extraction. |
| Outcome: | The proposed method acquires cross-document evidence and boosts performance in both closed and open environments. |
Copied to clipboard
| Challenge: | Existing benchmarks for multi-hop reasoning in biomedical domain are lacking . bioHopR provides benchmarks to evaluate multi-step reasoning in structured biomedic knowledge graphs . |
| Approach: | They propose a benchmark to evaluate multi-hop, multi-answer reasoning in biomedical knowledge graphs. |
| Outcome: | BioHopR evaluates multi-hop reasoning in biomedical knowledge graphs based on the PrimeKG model . it outperforms proprietary models and open-source biomedal models in 1-hop and 2-hop tasks . |
Copied to clipboard
| Challenge: | Standardized science questions require combining an average of 6 facts and as many as 16 facts to answer and explain. |
| Approach: | They propose to combine an average of 6 facts and as many as 16 facts to produce an answer for complex questions. |
| Outcome: | The proposed model is based on a corpus of 5,114 standardized science exam questions . it uses multi-fact explanations that combine science knowledge and world knowledge . |
Copied to clipboard
| Challenge: | Recent advances in context compression have failed to effectively utilize compressed representations for downstream tasks. |
| Approach: | They propose a holistic training paradigm that uses outcome-based RL to enable implicit expansion. |
| Outcome: | The proposed model outperforms previous models on NIAH, LongBench and multi-hop reasoning. |
Copied to clipboard
| Challenge: | Existing methods for post-training model editing suffer from overfitting and catastrophic forgetting. |
| Approach: | They propose a framework that leverages hyperbolic geometry and graph neural networks for precise and stable model edits. |
| Outcome: | Experiments on CounterFact, CounterFACT+, and MQuAKE with GPT2-XL and GPT-J show that HYPE significantly enhances edit stability, factual accuracy, and multi-hop reasoning. |
Copied to clipboard
| Challenge: | Existing multi-hop question answering models focus on multi-level reasoning across multiple documents or paragraphs. |
| Approach: | They propose a hierarchical graph network that aggregates clues from scattered texts . they use a set of contextual encoders to initialize nodes on different levels of granularity . |
| Outcome: | The proposed model outperforms existing multi-hop QA approaches on the HotpotQA benchmark. |
Copied to clipboard
| Challenge: | Existing attacks exploit leakage of retrieved subgraphs, leaving the security implications of structured knowledge representations unexplored. |
| Approach: | They propose a framework that leverages a novelty-guided exploration–exploitation strategy and external graph memory modules to extract a latent entity–relation graph. |
| Outcome: | The proposed framework outperforms baselines on medical, agriculture, and literary datasets under identical query budgets while maintaining high precision. |
Copied to clipboard
| Challenge: | Current RAG system retrieves evidence from knowledge graphs and text documents but has limitations in multi-hop reasoning, multi-entity questions, and source verification. |
| Approach: | They propose a training-free framework that unifies graph topology, document semantics, and source reliability to support deep, faithful reasoning in large language models. |
| Outcome: | The proposed framework outperforms the current hybrid model-based model-driven system by 20.3% and 30.1% on seven benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods rely on entity vector matching, but the purpose of the question is abstract and difficult to match with specific entities. Existing approaches rely only on entity-vector matching, and there is a problem with multi-hop reasoning. |
| Approach: | They propose a framework that constructs reasoning paths from purposes back to conditions using the KG ontology. |
| Outcome: | Experiments on the WebQSP and CWQ datasets show that ORT significantly improves the capability of large language models in knowledge graph question answering tasks (KGQA). |
Copied to clipboard
| Challenge: | Large language models suffer from factual inaccuracies in knowledge-intensive domains. |
| Approach: | They propose a question-guided KBQA framework that iteratively decomposes complex queries into simpler sub-questions and integrates a Graph Neural Network (GNN) to look ahead and incorporate 2-hop neighbor information at each reasoning step. |
| Outcome: | The proposed framework improves on four benchmark datasets and four LLMs. |
Copied to clipboard
| Challenge: | Vision-language models have shown impressive capabilities in perceptual tasks . however, they degrade in complex multi-hop reasoning under multi-player game settings . |
| Approach: | They propose a multi-agent framework for evaluating and synthesizing role-driven game scripts . they use curated and synthetic datasets to model uncertainty and deception . |
| Outcome: | The proposed model significantly boosts the performance of VLMs in narrative reasoning and hidden fact extraction under uncertain, adversarial, and socially complex conditions. |
Copied to clipboard
| Challenge: | Existing work on Temporal Question Answering (TQA) has focused on questions anchored to specific timestamps or events. |
| Approach: | They introduce a benchmark to address present-anchored temporal QA (PATQA) which includes single and multi-hop temporal questions. |
| Outcome: | The proposed model can be automatically refreshed by re-running SPARQL queries on a knowledge graph. |
Copied to clipboard
| Challenge: | Existing approaches to temporal knowledge graph question answering struggle with multi-hop reasoning and implicit temporal constraints. |
| Approach: | They propose a temporal tool-based API capable of transforming implicit temporal cues into executable operations and supervised fine-tuning teaches the model to interweave chain-of-thought reasoning with think-then-tool usage. |
| Outcome: | The proposed framework outperforms existing methods on three challenging questions. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated potential reasoning capabilities through prompt design, such as the Chain of Thought (CoT). |
| Approach: | They propose a new reasoning approach that predicts key entities which work as important “anchors” and employs a ranking algorithm to ensure the logical sequence of the predicted answers. |
| Outcome: | The proposed approach outperforms existing methods in multi-hop question reasoning and provides more accurate reasoning results in multihop question answering tasks. |
Copied to clipboard
| Challenge: | Existing methods for knowledge editing in Large Language Models face difficulties with multi-hop questions that require accurate fact identification and sequential logical reasoning. |
| Approach: | They propose a method that merges explicit knowledge representations of Knowledge Graphs with the linguistic flexibility of Large Language Models to convert free-form language into structured queries and fact triples. |
| Outcome: | The proposed method significantly surpasses state-of-the-art knowledge editing methods in the multi-hop question answering benchmark, MQuAKE. |
Copied to clipboard
| Challenge: | Multi-hop Question Answering (MHQA) adds layers of complexity to question answering tasks. |
| Approach: | They explore how LMs respond to multi-hop questions by permuting search results under various configurations. |
| Outcome: | The proposed model outperforms decoder-only models in MHQA tasks despite being significantly smaller in size . |
Copied to clipboard
| Challenge: | Existing retrieval augmented language models often overlook effective alignment with human preferences. |
| Approach: | They propose a benchmark to evaluate RMs in retrieval augmented language models . they incorporate 18 RAG subsets, six retrievers, and 24 RALMs to increase diversity . |
| Outcome: | The proposed benchmark combines 18 RAG subsets, six retrievers, and 24 RALMs to increase diversity of data sources. |
Copied to clipboard
| Challenge: | Existing graph RAGs decouple retrieval and reasoning processes, preventing adaptability . existing graph Raggings depend heavily on ground-truth entities, which are often unavailable in open-domain settings. |
| Approach: | They propose a graph retriever that is trained end-to-end with large-scale graphs . structure and semantic features are encoded via soft tokens and the verbalized graph . |
| Outcome: | The proposed approach improves the performance of large-scale graph retrieval models by grounding it with external knowledge. |
Copied to clipboard
| Challenge: | a rigorous and detailed comparison of the two spaces for multi-hop reasoning is lacking. |
| Approach: | They compare the capacity of hyperbolic space versus Euclidean space in multi-hop reasoning . they use an encoder-decoder model to integrate hyperbolical representations with a knowledge graph . |
| Outcome: | The proposed model outperforms the Euclidean space in multi-hop reasoning. |
Copied to clipboard
| Challenge: | Current large language models (LLMs) have shown a powerful ability of reasoning in understanding question requirement, retrieving the supporting fact and generating a precise answer. |
| Approach: | They propose a brain-inspired Theta-Gamma hierarchical oscillatory reasoning framework which decouples attention between global planning and local retrieval. |
| Outcome: | Extensive comparative experiments and specific validation experiments on multi-hop QA benchmarks show that THOR improves answer accuracy and robustness while mitigating limitations. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have facilitated the development of Multimodal LLMs. |
| Approach: | They propose a causal framework to interpret unimodal biases in visual question answering problems and a framework to integrate information from different modalities and mitigate biase. |
| Outcome: | The proposed framework analyzes visual question answering (VQA) problems to assess their impact on predictions. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have excellent performance in evaluation benchmarks, but struggle in complex reasoning tasks. |
| Approach: | They propose a tool-augmented chain-of-thought reasoning framework for chat-based LLMs . they model chain- of-thoughting reasoning as multi-turn conversations to utilize tools . |
| Outcome: | The proposed framework can outperform state-of-the-art models on complex reasoning tasks. |
Copied to clipboard
| Challenge: | Existing methods rely on supervision for both answers and rationales, but they have limited capacities in modeling interactions between sentences, let alone reasoning across multiple documents. |
| Approach: | They propose a principled, probabilistic approach for training explainable multi-hop question answering systems without rationale supervision. |
| Outcome: | The proposed method is more accurate at selecting rationales than previous methods while maintaining similar accuracy in predicting answers. |
Copied to clipboard
| Challenge: | Existing methods for Graph-based retrieval-augmented generation rely on implicit semantic relevance propagation. |
| Approach: | They propose a semantic-aware retrieval framework that improves both semantic recall and explicit reasoning. |
| Outcome: | Extensive experiments show that FlowRAG improves both semantic recall and explicit reasoning. |
Copied to clipboard
| Challenge: | Existing benchmarks address single tables or non-visual data, leaving a critical gap . MTabVQA comprises 3,745 complex question-answer pairs . |
| Approach: | They propose a benchmark specifically designed for multi-tabular visual question answering that measures the ability to parse diverse table images and correlate information across them. |
| Outcome: | The proposed benchmarks show that fine-tuning VLMs with MTabVQA-Instruct significantly improves their reasoning abilities. |
Copied to clipboard
| Challenge: | Current retrieval-augmented generation methods struggle with complex multi-hop reasoning, relying on unstructured semantic matching that lacks the logical structure needed to systematically guide retrieval. |
| Approach: | They propose a framework that elevates retrieval to structured, program-guided reasoning by combining three stages of program-type selection and evidence accumulation. |
| Outcome: | Evaluated on five benchmarks including HotPotQA, 2WikiMultihopQA, ARC-Challenge, MMLU-Pro, and MedQA with various LLMs, PROGRAM achieves state-of-the-art performance with up to 24% relative improvement on HotPtQA and 13.2% on MedQA over strong baselines including FLARE, ProbTree and Self-RAG. |
Copied to clipboard
| Challenge: | Recent work has demonstrated unprecedented capabilities in sophisticated linguistic comprehension and generative tasks. |
| Approach: | They propose a framework for search LLMs that trains with step-wise proximal policy optimization method to improve QA performance. |
| Outcome: | The proposed framework outperforms global-reward benchmarks on multi-hop QA with a stepwise proximal policy optimization method and richer and more detailed intermediate search rewards and token-level process supervision. |
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models lack data contamination and complex queries . financial cross-modal multi-hop reasoning is difficult to evaluate and requires precise cross-module reasoning . |
| Approach: | They propose a benchmark to analyze the reasoning capabilities of multimodal large language models. |
| Outcome: | The proposed model is categorized into three difficulty levels—easy, medium, and hard—for step-by-step evaluation. |
Copied to clipboard
| Challenge: | Existing models struggle to address two major bottlenecks in Document-level Relation Extraction: extreme class imbalance and complexity of multi-hop reasoning. |
| Approach: | They propose a method that decouples the extraction space into dense and sparse scenarios. |
| Outcome: | The proposed approach yields consistent improvements over various backbone models and achieves advanced performance compared to existing enhancement methods. |
Copied to clipboard
| Challenge: | a prior graph-based approach to global sensemaking lacks retrieval mechanisms, topic specificity, and incurs high inference costs. |
| Approach: | They propose a RetrievalEnhanced, Topic-Augmented Graph framework that retrieves relevant summaries from a topic. |
| Outcome: | The proposed framework improves response quality while significantly reducing inference time compared to the baseline. |
Copied to clipboard
| Challenge: | Existing KG-RAG systems collapse all reasoning hops into a single representation, flat embedding space, suppressing this implicit structure and causing noisy or drifted path exploration. |
| Approach: | They propose a symmetric multi-view framework that decouples queries and KGs into aligned, head-specific retrieval spaces. |
| Outcome: | The proposed framework achieves state-of-the-art retrieval and QA performance on WebQSP and CWQ, and significantly reduces hallucination. |
Copied to clipboard
| Challenge: | Existing lightweight approaches to retrieval-augmented generation fail to capture latent semantic connections between disjoint entities. |
| Approach: | They propose a lightweight RAG framework that constructs a hypergraph capturing both structure and semantic relationships using a hybrid structural-semantic retrieval mechanism. |
| Outcome: | EHRAG outperforms state-of-the-art methods on four datasets while maintaining zero token consumption. |
Copied to clipboard
| Challenge: | Existing knowledge editing approaches directly edit model context without isolating target knowledge from the reasoning path of model inference, resulting in unreliable and low-quality outputs, especially in multi-hop tasks. |
| Approach: | They propose a framework that separates model reasoning from knowledge editing and propose 'DecKER' that allows users to modify specific factual associations without retraining the entire model. |
| Outcome: | The proposed framework significantly improves multi-hop reasoning performance by mitigating knowledge conflicts and preserving reasoning integrity. |
Copied to clipboard
| Challenge: | Vision-Language Models have shown impressive capabilities and notable failures in data visualization understanding tasks. |
| Approach: | They propose a benchmark to analyze how specific properties within a visualization type affect VLM performance. |
| Outcome: | The proposed benchmark examines how specific properties affect VLM performance . it shows that models exhibit steep drops on multi-hop reasoning and extraction errors increase with edge density . |
Copied to clipboard
| Challenge: | Existing methods for linking knowledge graphs are incomplete and rely on Euclidean embeddings . a hyperbolic GNN framework embeds recursive learning trees in hyperbolical space . |
| Approach: | They propose a hyperbolic GNN framework that embeds recursive learning trees in hyperbolical space and generates query-specific embeddings. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on multiple benchmark datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on well-structured tables and fail to reflect irregular structures and complex reasoning commonly encountered in real-world scenarios. |
| Approach: | They propose a benchmark to evaluate TableQA under complex reasoning and irregular table conditions. |
| Outcome: | The proposed framework improves generalization and realism of large language models under complex and irregular table conditions. |
Copied to clipboard
| Challenge: | Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing LLMs’ short-context reasoning but falters in long-contemporal scenarios requiring precise grounding and multi-hop reasoning. |
| Approach: | They propose a framework that constructs high-difficulty, multi-hop long-context QA pairs with inherent reasoning chains to overcome this bottleneck. |
| Outcome: | The proposed framework outperforms RLVR baselines and matches frontier LLMs while using far fewer parameters. |
Copied to clipboard
| Challenge: | Existing retrieval-augmented generation (RAG) methods fail to provide deep, relational understanding of scientific literature. |
| Approach: | They propose a graph-grounded reasoning framework for structured scientific evaluation that uses multi-hop reasoning to iteratively construct contextual graphs and generate structured critiques. |
| Outcome: | The proposed framework reduces evaluation error by over 30% compared to baselines and allows smaller models to outperform larger models. |
Copied to clipboard
| Challenge: | Current large language models struggle to answer questions that span tens of thousands of tokens. |
| Approach: | They evaluate 1–4 hop QA over 64k–128k-token excerpts from 83 novels . they find consistent accuracy drops with increased hops and context length . |
| Outcome: | The novelhopqa benchmark evaluates 1–4 hop QA over 64k–128k-token excerpts from 83 public-domain novels. |
Copied to clipboard
| Challenge: | Existing knowledge editing techniques show limitations when applied to multi-hop reasoning . residual single-hop knowledge causes edited models to revert to original answers . |
| Approach: | They propose a knowledge editing method that incorporates a Knowledge Erasure mechanism for Large language model Editing (KELE) they propose an erasure function for residual knowledge and an injection function for new knowledge . |
| Outcome: | The proposed method significantly improves multi-hop reasoning capability of edited models. |
Copied to clipboard
| Challenge: | GraphCheck is a framework for fact-checking complex claims that require multi-hop reasoning . Graphcheck excels in complex scenarios, but may be unnecessarily elaborate for simpler claims . |
| Approach: | They propose a framework that transforms claims into entity-relationship graphs for fact-checking . DP-GraphCheck employs a lightweight strategy selector to choose between direct prompting and GraphCheck adaptively. |
| Outcome: | The proposed framework outperforms existing methods in verification accuracy while achieving strong computational efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability, but they may not reflect the model’s truereasoning faithfully. |
| Approach: | They propose a testbed framework for evaluating faithfulness metrics for natural language explanations using diagnosticity and model-editing methods. |
| Outcome: | The proposed framework evaluates faithfulness metrics for natural language explanations on four tasks including fact-checking, analogy, object counting, and multi-hop reasoning. |
Copied to clipboard
| Challenge: | Existing graph-based or hybrid systems lack the ability to integrate supplementary evidence as reasoning unfolds. |
| Approach: | They propose a framework that integrates non-parametric knowledge into Large Language Models . they use a RL-based framework to optimize the entire generation process via RL . |
| Outcome: | The proposed framework outperforms existing RAG frameworks in five question answering benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for table-text retrieval are limited due to the need to bridge structured tables and unstructured passages. |
| Approach: | They propose a table-text retrieval system that combines the strengths of both approaches . they propose bipartite subgraph retrieval and query-relevant node expansion . |
| Outcome: | The proposed method outperforms state-of-the-art models with a 42.6% and 39.9% improvement on the OTT-QA benchmark. |
Copied to clipboard
| Challenge: | Existing prompt compression methods are designed for single-turn queries and fail to capture interdependent reasoning steps. |
| Approach: | They propose a unified, training-free prompt compression framework that integrates multi-hop reasoning within an iterative compression loop. |
| Outcome: | Experiments on MusiQue, 2WikiMultiHopQA, and HotpotQA show that iterCOMP achieves significant improvements in Exact Match and F1 scores while reducing the token budget. |
Copied to clipboard
| Challenge: | Existing methods for integrating knowledge graphs with large language models lack continuous learning capabilities. |
| Approach: | They propose an agent framework with a dynamic, evolvable memory mechanism specifically designed for KG reasoning. |
| Outcome: | EvoMemKG achieves state-of-the-art performance without training or tools . it achieves improvements of up to 20% over baseline on multi-hop queries . |
Copied to clipboard
| Challenge: | Current methods for evaluating LLMs’ veracity are limited by the need for extensive human labor, test data contamination, or limited scope, hindering efficient and effective exposure of errors. |
| Approach: | They propose a framework that extracts fact triplets to generate diverse question types using rule-based natural language processing techniques. |
| Outcome: | The proposed framework can trigger factual errors in up to 55% of questions in large LLMs while maintaining coverage of questions. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) enables large language models to incorporate external knowledge at inference. |
| Approach: | They propose a hierarchical framework that organizes triplets into subtopics and topics to enhance connectivity and integrate dispersed information. |
| Outcome: | Experiments on abstractive and specific QA benchmarks show that TH-RAG outperforms strong baselines in accuracy and robustness while remaining efficient. |
Copied to clipboard
| Challenge: | Existing knowledge-based VQA benchmarks focus on coarse-grained categories and simple reasoning over single entities. |
| Approach: | They propose a knowledge-based Visual Question Answering benchmark to enhance multimodality evaluation. |
| Outcome: | The proposed benchmark improves evaluation of multimodal large language models in fine-grained multimodal entity understanding and complex multihop reasoning. |
Copied to clipboard
| Challenge: | Existing paradigms for multi-hop reasoning suffer from high construction costs and limited adaptability to dynamic knowledge. |
| Approach: | They propose a symbolic reasoning framework for multi-hop question answering that integrates the advantages of both paradigms by dynamically generating sub-questions, performing information retrieval and symbolic encoding based on an on-the-fly graph and using a symbol verifier to validate intermediate reasoning steps. |
| Outcome: | The proposed framework significantly improves accuracy and robustness on multiple multi-hop benchmarks and a medical dataset. |
Copied to clipboard
| Challenge: | Existing studies have identified a position bias in Large Language Models that causes them to overlook information at certain positions. |
| Approach: | They propose a semantic probe to disentangle position bias in Large Language Models . they propose MFAI to steer attention towards selected positions . |
| Outcome: | The proposed model can locate and integrate information at certain positions even in noisy, long-context settings. |
Copied to clipboard
| Challenge: | Existing graph-based methods for enhancing Large Language Models (LLMs) with external knowledge are focusing on local relationships, resulting in suboptimal performance for tasks that require global context. |
| Approach: | They propose a "panorama"-guided paradigm that integrates a light yet comprehensive "panoramic" of the corpus to guide all stages of the retrieval process. |
| Outcome: | The proposed paradigm performs well across five datasets and a variety of tasks. |
Copied to clipboard
| Challenge: | Existing methods for understanding video over long periods of time are limited . eGAgent system provides tools for structured search and reasoning over entity scene graphs . |
| Approach: | They propose a system that can interpret and recall video over days or weeks . they use entity scene graphs to equip a planning agent with tools for structured search and reasoning . |
| Outcome: | The proposed method achieves state-of-the-art performance on EgoLifeQA and Video-MME-long datasets. |